Genome Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Genome Medicine's content profile, based on 183 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.
Uria-Regojo, G.; Fernandez-Caballero, L.; Lopez-Alcojor, A.; Lopez-Lopez, L.; Benitez, Y.; Rodilla, C.; Avila Fernandez, A.; Trujillo-Tiebas, M. J.; Osorio, A.; Corton, M.; Almoguera, B.; Ayuso, C.; Minguez, P.
Show abstract
Rare diseases (RDs) remain a major diagnostic challenge. Genetic and phenotypic heterogeneity, incomplete knowledge of disease mechanisms, and limitations in variant clinical interpretation leave many patients without a molecular diagnosis. Meanwhile, the growing volume of genomic data generated in clinical practice offers an opportunity to develop data-driven methodologies for exploring disease mechanisms and improving the reanalysis of unsolved cases. We aggregated real-world genomic data from 11,084 unrelated patients with suspected RD. Patients were clinically classified into 122 diseases. We built a multi-disease genomic variant frequency database (FJD-DB), which enabled the development of variant and gene-disease association scores by means of case-control subcohort comparisons across 32 disease groups. Functional enrichment analyses were then used to highlight disease-associated protein domains, pathways, biological processes, and phenotypes. Finally, the resulting knowledge was integrated into a data-driven framework for the guided reanalysis of unsolved RD patients applied to Inherited Retinal Dystrophies (IRD) patients as first use case. FJD-DB contained more than 45 million unique variants, including ~185,000 potentially pathogenic variants. Disease-specific analyses identified disease-associated pathogenic variants and highlighted both established and candidate disease genes. We detected 179 significantly enriched protein domains across 23 diseases, 124 Human Phenotype Ontology terms across 13 diseases, 79 Reactome pathways across 10 diseases, and 72 Gene Ontology biological processes across 8 diseases, revealing highly disease-specific functional signatures. Integration of disease-specific variant, gene, and functional association signals enabled the development of a data-driven framework for guided reanalysis of unsolved RD cases. Applied to more than 1,100 unsolved IRD cases, the framework generated clinically relevant findings in 26 patients, including four molecular diagnoses, seven candidate diagnoses, and 15 cases upgraded from non-informative findings to variants of uncertain significance. Aggregated real-world genomic data can be leveraged to identify disease-associated molecular signals generating novel biological hypotheses. A unified analytical framework provides a scalable strategy for knowledge discovery and guided reanalysis, facilitating the identification of overlooked and potentially novel genetic causes of RDs.
Ou, Y.; Zhao, S.; Weng, J.; Long, Y.; Li, H.; Hou, Y.; Xiong, Q.; Liu, S.; Wang, Z.; Xu, Y.; Pan, H.; Zhang, H.; Sun, L.
Show abstract
Gut microbial structural variation captures strain-level genomic diversity beyond species abundance, yet its contribution to rheumatoid arthritis (RA) remains unclear. Integrating gut metagenomic datasets from four independent cohorts (n = 491), we identified reproducible RA-associated structural variants (SVs), most of which occurred in species without differential abundance. Incorporating SVs into machine learning models consistently improved disease discrimination across independent validation cohorts. Functional analyses identified a core deletion SV in Agathobacter rectalis that removes an XRE-family transcriptional regulator. Motif discovery and 3D structural modeling demonstrated sequence-specific binding of the regulator to the promoter of a short-chain fatty acid biosynthetic gene, supporting a strain-level regulatory mechanism. Together, these findings establish microbial structural variation as a complementary functional layer beyond taxonomy for RA discrimination and mechanistic interpretation.
Rota Negroni, M.; Billato, I.; Romualdi, C.
Show abstract
Copy number alterations (CNAs) are major contributors to genomic instability in cancer, and copy number signatures (CNS) provide a compact representation of the processes shaping CNA landscapes. However, the relationships among existing CNS frameworks and their predictability from molecular data other than whole-genome sequencing remain unclear. Here, we compare three major CNS compendia across more than 5,800 TCGA cancer samples, evaluating their overlap, complementarity, biological relevance, and prognostic associations. Individual signatures showed limited cross-study concordance, whereas signature-derived clusters identified biologically distinct patient groups, including favorable-outcome clusters observed across all frameworks. Using gene expression, DNA methylation, somatic mutation features, age, and tumor purity, XGBoost models predicted cluster membership with framework-dependent performance, achieving high F1-scores for the Drews and Steele compendia but limited performance for Tao. Feature importance analysis highlighted expression-driven predictors and pathways linked to genomic instability. These findings show that current CNS frameworks capture complementary rather than interchangeable dimensions of tumor genome instability and suggest that multi-omic profiles can extend signature-based stratification to cohorts without whole-genome sequencing.
Bermudez-Guzman, L.; Ramos-Esquivel, A.; Alpizar-Alpizar, W.
Show abstract
Early- and average-onset colorectal cancer (CRC) are separated at age 50, but whether this defines a biological threshold remains unclear. To clarify this, we identified molecular profiles in nine harmonised cBioPortal CRC cohorts (4,609 patients) by fitting Bernoulli mixture models to 31 repair-state, genomic-burden and gene-alteration features, excluding age, sex and tumour site, and compared their prevalence using <50/[≥]50 and decade-resolved groups. Four profiles captured conventional/CIN-like (P1), intermediate MSS (P2), KRAS/PI3K/APC-rich (P3) and hypermutated/MSI-high (P4) states along a left-to-right gradient. Although molecular identities remained stable, profile prevalence followed non-linear P1/P4 and opposing linear P2/P3 age trajectories. Profile-prevalence patterns did not track chronological proximity: profile composition at 30-39 differed from 50-59 but not clearly from 60-69. The age-50 threshold captured only 17.8% of decade-resolved deviance, whereas the optimal age-70 cut-off retained only 51.3%. Validation in 2,579 non-overlapping MSK-IMPACT patients (2,476 age-evaluable) reproduced molecular-feature patterns (r=0.97-0.98), age trajectories (r=0.92) and limited binary-threshold performance: age 50 and the optimal age-66 cut-off retained 12.7% and 45.1%, respectively. Thus, age reorganizes the prevalence of shared CRC states rather than defining a biological threshold at age 50.
Kulkarni, R.; Sengupta, A.; Kumar, R.
Show abstract
Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity, which complicates tumour stratification and limits the ability of mutation-centric models to capture tumour behaviour and predict patient outcomes. This study investigates whether coordinated transcriptomic programs can provide a systems-level representation of tumour states. Bulk RNA-sequencing data from the TCGA-LUAD cohort were analysed to reconstruct pathway-level transcriptomic organisation using a stability-optimised network framework (SPARK). This analysis identified eight transcriptomic modules representing coordinated biological processes active across tumours. Module activity scores were subsequently used to derive a composite Transcriptomic Risk Score through elastic-net Cox proportional hazards modelling. The resulting risk score showed a significant association with overall survival in the discovery cohort and improved prognostic discrimination beyond clinical variables. An independent evaluation in the CPTAC-LUAD cohort confirmed the prognostic signal and preserved risk stratification across patient groups. Unsupervised clustering of module activity further revealed three transcriptomic patient groups characterised by distinct biological programs, genomic alteration patterns, and survival outcomes. Single-cell analysis also demonstrated that the identified transcriptomic modules reflect coordinated organisation of the tumour-immune-stromal ecosystem across cellular compartments. Together, these findings suggest that LUAD heterogeneity can be organised into coordinated transcriptomic programs with measurable clinical relevance, providing a systems-level framework for representing tumour molecular states.
Nakagawa, H.; Kamatani, T.; Ishibashi, N.; Aoyama, S.; Morioka, M.; Miya, F.; Ikeda, S.
Show abstract
Comprehensive genomic profiling (CGP) supports precision medicine in cancer care, but accurate assessment of missense variant pathogenicity, especially for variants without established consensus, remains challenging. Various computational tools have been developed for variant functional prediction, but most current tools rely solely on variant-level features and do not capture the clinical context of individual patients. To address this limitation, we developed MARiO (Missense Alteration Risk for Oncogenicity), a machine-learning model that integrates variant-level features and patient-level clinical and genomic contexts to effectively predict the pathogenicity of missense variants in cancer. We collected a total of 10,642 missense variants from 1271 patients, and evaluated candidate features for their association with variant pathogenicity, identifying informative features including in silico functional predictions, population allele frequency, variant allele frequency, and tumor mutational burden. Using these selected features, MARiO was developed with extreme gradient boosting. The model integrates multiple in silico prediction tools and patient-specific genomic contexts while accommodating missing values frequently observed in real-world CGP datasets. MARiO outperformed existing tools, achieving an area under the receiver operating characteristic curve of 0.942. The model demonstrated strong generalizability across multiple external datasets and showed consistency with real-world molecular treatment proposals. MARiO offers a robust and clinically relevant approach for missense variant pathogenicity assessment by integrating variant- and patient-level features and serves as a valuable tool to support clinical decision-making.
Cilleros Portet, A.; Gonzalez-Moro, I.; Saddiki, H.; Mari, S.; Everson, T.; Hernangomez-Laderas, A.; Broseus, L.; Tost, J.; Cosin-Tomas, M.; Groleau, M.; Czamara, D.; Lozano, M.; Hao, K.; Tuhkanen, J.; Fallin, M. D.; Schmidt, R. J.; Breeze, C. E.; Deleuze, J.-F.; Aguilar-Lacasana, S.; Jacques, P.-E.; Hytti, S.; Irizar, A.; Lahti-Pulkkinen, M.; Bakulski, K. M.; Dou, J.; Lahti, J.; Vrijheid, M.; Raikkonen, K.; Hivert, M.-F.; Sunyer, J.; Heude, B.; lepeule, j.; London, S. J.; Chen, J.; Bustamante, M.; Marsit, C.; Bilbao Catala, J. R.; Lesseur, C.; Fernandez-Jimenez, N.
Show abstract
The Developmental Origins of Health and Disease (DOHaD) hypothesis proposes that the perinatal environment shapes susceptibility to complex traits across life [1]. The placenta, a transient organ mediating maternal-fetal exchange, plays a central role in this process and has emerged as a key molecular archive in utero [2-4]. Placental DNA methylation (DNAm) is a unique mediator between prenatal exposures, fetal genetics and later-life outcomes [5-9]. DNAm quantitative trait loci (mQTL) have helped disentangling causal mechanisms underlying GWAS loci for complex diseases [10-15]. Despite growing evidence that placental genomic regulation has broad and profound effects on the developmental programming of early- and later-life health outcomes [17], existing placental studies remain limited in scale and largely focused on growth- and neuro-related traits [12-16]. Here, we construct a high-resolution placental mQTL resource and systematically investigate how placental DNAm relates to early- and later-life traits, and to shared vulnerability and complex interactions among them.
Tallman, D.; Striker, S.; Byappanahalli, A. M.; Stockard, S.; Jenison, J.; Collier, K. A.; Blige, E.; Vater, M.; Stover, D. G.
Show abstract
BackgroundCopy number aberrations (CNAs) are gains and losses of large genomic segments present across most cancer types and are a hallmark of cancer genomic alterations. However, the processes underlying CNAs and characteristic patterns of CNAs are poorly understood. Bioinformatic advances have identified underlying single nucleotide variant (SNV) mutational signatures resulting from distinct mutational processes, yet development of algorithms able to uncover similar signatures for CNAs remains less advanced. MethodsUsing segmented data files from DNA sequencing, six copy number features are extracted for signature determination: segment size, breakpoints per 10 megabases, copy number oscillation events, average changepoint size, average copy number, and breakpoints per chromosome arm, along with ploidy. Mixed model approaches and non-negative matrix factorization (NMF) are utilized to derive CNA signatures across cancer types. The full methodology was packaged in a robust R package, termed CNSigs that is publicly available. ResultsTo verify the reproducibility of the signatures, we derived five signatures from two independent breast cancer datasets (total n>3000), demonstrating high accuracy (average cosine similarity = 0.89). Pan-cancer application of CNSigs in the TCGA dataset resulted in derivation of 13 pan-cancer signatures which were significantly associated with disease-specific survival. Benchmarking CNSigs to two other CNA signature approaches within TCGA demonstrated non-overlapping signatures and favorable compute speed for CNSigs. We evaluated n=24 pairs of tumor and circulating tumor DNA (ctDNA) acquired at the same time and demonstrated that CNSigs are detectable and reproducible via ctDNA, with significant association of CNSig11 with metastatic triple-negative breast cancer progression-free survival for taxane but not platinum or capecitabine chemotherapy. CNSigs association with immunophenotype was evaluated in low-grade glioma (LGG) and CNSig 3 was found to be highly prognostic for LGG yet complementary to immune features. ConclusionsThe CNSigs R package allows researchers to easily analyze their own samples to derive copy number signatures and evaluate clinical associations. We demonstrate potential application in ctDNA and association with treatment response. The development of this package allows further investigation of underlying processes that may be responsible for these CNA fingerprints.
Jiang, K.; Jarvis, J. N.
Show abstract
While progress has been made in identifying the true risk-driving single nucleotide polymorphisms (SNPS) on juvenile idiopathic arthritis (JIA) risk haplotypes, the affected cells and target genes largely remain unknown. We used data from a previously published massively parallel reporter assay (MPRA) to query human data in the Database of Immune Cell eQTLs (DICE) and the Gene-Tissue Expression (GTEx) database to identify affected cells and target genes of MPRA-identified SNPs in immune cells and relevant tissues. SNPs identified on MPRA were associated with gene expression levels in a broad range of immune cells in the DICE database, including CD4+ and CD8+ T lymphocytes, monocytes, NK cells, and B cells. MPRA-identified SNPs showed strong associations with gene expression in GTEx whole blood, spleen, and/or EBV-stimulated lymphocytes. Our data show the efficacy of combining MPRA and using human cells/tissue expression data to elucidate complex mechanisms driving genetic risk for JIA.
Nielsen, M. C.; Mentzel, C. M. J.; Stoltze, U. K.; Hagen, C. M.; Baekvad-Hansen, M.; Byrjalsen, A.; Sunde, L.; Lundquist, A. A.; Lund, A. M.; Tfelt-Hansen, J.; Masmas, T.; Soerensen, E.; Pedersen, O. B. V.; Erikstrup, C.; Ostrowski, S. R.; DBDS Genomic Consortium, ; Hjalgrim, H.; Nyegaard, M.; Schmiegelow, K.; Hansen, T. v. O.; Wadt, K.; Bybjerg-Grauholm, J.; Rasmussen, S.
Show abstract
Genetic screening for rare pathogenic variants facilitates early detection and prevention of disease manifestations in medically actionable disorders, but sequencing costs limit widespread use. We introduce DoBSeq, a low-cost, high-throughput screening framework for detecting rare, single-nucleotide variants and indels. The framework includes: extraction of DNA from dried blood spots used in neonatal screening, automation of two-dimensional DNA pooling and library preparation, high-depth targeted sequencing using a 582-gene custom panel, and a probabilistic model to assign rare pathogenic variants to individuals. Benchmarked against whole-genome sequencing across 582 genes in a batch of 576 individuals, the framework detected 95% of all variants and recovered all clinically relevant pathogenic single-nucleotide variants in American College of Medical Genetics and Genomics (ACMG) actionable genes. Applied to 2304 anonymised blood donors, it yielded variant frequencies consistent with existing population estimates. At a sample cost of 29 USD, including 11 USD running costs, this framework provides a cost-efficient approach to population-level genetic screening.
Li, Q.; Xu, L.; Wang, J.; Li, C.; Wen, W.; Shu, X.; Yang, Y.; Shu, X.-o.; Cai, Q.; Long, J.; Singh, B.; Lau, K. S.; Yin, Z.; Casey, G.; Song, M.; Peters, U.; Zheng, W.; Guo, X.
Show abstract
Bulk tissue-based DNA methylation-wide (MWAS) and transcriptome-wide association studies (TWAS) have identified CpG sites and genes associated with colorectal cancer (CRC) risk, but do not account for cellular heterogeneity. To address this, we developed a deconvolution-informed framework to infer cell-type specific DNA methylation and gene expression profiles from bulk normal colon tissues using reference single-cell epigenomic and transcriptomic datasets. We performed cell-type specific MWAS (ctMWAS) using deconvoluted DNA methylation data from 293 normal colon samples and conducted cell-type specific TWAS (ctTWAS) using deconvoluted gene expression data from 707 normal colon samples. Genetically predicted methylation and expression models were integrated with CRC GWAS summary statistics (78,473 cases and 107,143 controls) to identify risk-associated CpG sites and genes. Through ctMWAS, ctTWAS, and colocalization analyses, we identified 178 significant cell-type-specific CpG sites in 106 loci and 68 risk genes in 40 loci, including 26 previously unreported loci. Through additional integrative methylation-gene analysis, we prioritized 132 candidate risk genes, the majority of which were supported by multi-omics evidence and stage-specific dysregulation across the adenoma-carcinoma and serrated-carcinoma progression pathways. Pathway enrichment analyses implicated pathways involved in DNA double-strand break repair, TP53 regulation, TGF-{beta} signaling, and innate immune responses. Among prioritized genes, 14 were identified as putative druggable targets linked to 90 FDA-approved or clinical-stage drugs. Experimental validation supports an oncogenic role for SF3A3. These findings demonstrate that deconvolution-informed integrative analyses enable cell-type-resolved identification of epigenetic and transcriptional mechanisms underlying CRC susceptibility and provide insights into disease biology, prevention, and therapeutic target discovery.
Multerer, K.; Atkinson, P.; Woods, L.; Tanigawa, Y.; Kellis, M.; Munkacsi, A.
Show abstract
Polygenic risk scores (PRS) assume additive SNP effects, yet genetic risk also arises from interactions between loci and environmental factors that contribute to broad-sense heritability. We developed an extended PRS (ePRS) framework for type 2 diabetes (T2D) that incorporates locus-by-locus non-additive effects beyond those captured by additive single-locus PRS or linkage disequilibrium (LD) tagging. These were modelled as cumulative burden (G+G; summed allele counts), statistical epistasis (GxG; allele count products), and gene-environment effects derived from cardiometabolic variables in electronic health records. Across 235,000 UK Biobank participants, five complementary ePRS models captured largely non-overlapping high-risk individuals, suggesting that a key to individual risk predictions comprise the inclusion of multiple interaction-driven biological components rather than a single signal. A composite score improved case detection beyond clinical predictors, including individuals within clinically normal ranges. These findings were generalized to celiac disease, with similar complementarity across models, with potential for clinical use pending prospective validation.
Farid, A. C.; Haldeman, S.; Otto, C.; DMello, A.; Tettelin, H.; Ratner, A. J.
Show abstract
Based on recent epidemiologic studies, Streptococcus agalactiae (Group B Streptococcus; GBS) sequence type (ST) 1010 is an emerging lineage now identified in multiple countries. We report the phylogenetic and genomic characteristics of a set of 55 GBS sequence type (ST) 1010 strains, as well as two newly described single-locus variants of ST1010. A core genome phylogeny suggests that ST1010 is closely related to both ST452 and the hypervirulent clonal complex (CC) 17 GBS lineage. Notably, we demonstrate that genes encoding two virulence determinants previously described as specific to CC17 GBS, the HvgA adhesin and the serine-rich repeat protein Srr2, are both present in ST1010 genomes. Srr2 is shared with members of ST452. High-level gentamicin resistance (HLGR) encoded on an IS256 mobile element, previously described in a small number of ST1010 isolates, is present in a distinct ST1010 subclade encompassing the majority of ST1010 isolates. The relationship between ST452 (serotype IV), ST1010 (serotype IV), and ST17 (serotype III) strains suggests that ST17 may have arisen from a serotype IV ancestor and later acquired the type III capsule locus. Taken together, these findings clarify the phylogenetic position of ST1010 and suggest sequential acquisition of virulence determinants and HLGR prior to its international emergence. IMPACT STATEMENTST1010 GBS has emerged internationally, with colonizing and invasive isolates described in the United States, Dominican Republic, Netherlands, and Italy. Using a core genome phylogeny and targeted detection of genomic regions, we demonstrate that ST1010 shares specific virulence determinants with the CC17 hypervirulent GBS lineage and that HLGR is confined to a specific numerically dominant subclade of ST1010. Our work spotlights the importance of future epidemiologic and genomic surveillance of ST1010 and related lineages. DATA SUMMARYPublicly available genomic data were used from three previously published studies (Laycock KM et al., McGee L et al., Khan UB et al.), as well as a set of newly sequenced GBS genomes from clinical strains originating in New York City (NYC). The corresponding accession numbers and detailed information for all strains are provided in the Table.
Manojlovic, V.; Gabbutt, C.; Shibata, D.; Noble, R. J.
Show abstract
Determining the nature of human tumour growth is challenging given the impracticality of obtaining detailed data across time. A promising solution is to examine DNA regions whose methylation states fluctuate on clinically relevant timescales, permitting their use as high-resolution lineage tracers. However, existing methods developed for analysing the fluctuating methylation loci of normal tissue and lymphoid cancers are inapplicable to large solid tumours. Here we introduce a mechanistic computational model that tracks the evolution of heritable methylation marks as a tumour grows from a single gland to a mass of many cubic centimetres, and a coupled ABC-SMC inference workflow to estimate tumour growth parameters from multi-region bulk methylation arrays. We applied this framework to data from multiple regions of 10 resected colorectal tumours, including 3 adenomas and 7 carcinomas of diverse sizes and clinical stages. By exploring alternative models, we show that intratumour diversity, in terms of methylation errors, stems more from tumour growth via gland fission than from cell turnover within glands. Moreover, the extent of intratumour diversity varies widely between patients, mainly because of eight-fold variation in gland fission rates but also due to differences in methylation and demethylation rates. Intergland divergence patterns are consistent with neutral evolution of colorectal tumours and a cancer stem cell fraction of approximately 1%. As well as helping to resolve the nature of colorectal cancer growth and evolution, our results provide proof of principle for a method that may be adapted to other types of solid tumour.
Demir, A. Y.; Yasar, E.
Show abstract
Integrated prognostic signatures combining ferroptosis, cuproptosis, and disulfidptosis are increasingly reported in oncology as advances in risk stratification, yet their added value over simpler pathway-specific or proliferation-related models remains unclear. Here, we developed an integrated regulated cell-death signature and evaluated it through an adversarial pan-cancer benchmark. Using the TCGA pan-cancer cohort comprising 9,808 tumours across 33 cancer types, we curated 118 genes associated with the three cell-death programmes, characterised inter-pathway crosstalk, and derived a 26-gene LASSO-Cox risk signature. The model showed reproducible prognostic performance across cancers, with a pan-cancer concordance index of 0.573 (95% CI, 0.552-0.594), and was independently validated in METABRIC and CGGA cohorts, remaining significant after adjustment for standard clinical variables. However, benchmarking revealed that the integrated signature, although superior to size-matched random gene sets (empirical p < 0.001), did not outperform a ferroptosis-only model (DeLong p = 0.81), indicating no measurable gain from pathway integration. Moreover, much of the prognostic signal reflected tumour proliferation rather than regulated cell death. After adjustment for the proliferation meta-signature (meta-PCNA), ferroptosis performance declined from 0.573 to 0.504, while the integrated model decreased to 0.554. High-risk tumours were more sensitive to anti-proliferative drugs, and the risk score was most strongly associated with E2F, MYC, and G2M target programmes. The signature stratified prognosis but did not predict immune-checkpoint blockade response in IMvigor210 (AUC {approx} 0.50). Importantly, the underlying biology was not merely a modelling artefact. Signature genes showed concordance with protein abundance in CPTAC cohorts, and the three cell-death programmes co-varied within individual malignant cells, with correlations ranging from {rho} = 0.46 to 0.66. Overall, our findings indicate that integrated multi-death signatures are reproducible and biologically grounded, yet prognostically redundant and substantially confounded by proliferation. This study provides a cautionary benchmark for the rapidly expanding use of composite regulated cell-death signatures in cancer prognosis.
Allen, S.; Rowlands, C. F.; Kuzbari, Z.; Garrett, A.; Durkie, M.; Burghel, G. J.; Robinson, R.; Callaway, A.; Field, J.; Frugtniet, B.; Palmer-Smith, S.; Grant, J.; Pagan, J.; Johnston, E.; McDevitt, T.; Hughes, L.; Yarram-Smith, L.; Logan, P.; Reed, L.; Snape, K.; McVeigh, T.; Hanson, H.; Roth, F. P.; Starita, L. M.; Fowler, D. M.; Villani, R.; Spurdle, A. B.; Adams, D. J.; Findlay, G.; Turnbull, C.; Cancer Variant Interpretation Group UK (CanVIG-UK),
Show abstract
Background: Clear guidance is lacking regarding how 'truthset' variants should be used for clinical validation of functional assays, namely determining the allocatable evidence points (EPs) towards clinical classification. It is argued that assays should be validated using truthsets of missense variants, as this is the variant type for which classification is most impacted by functional data. EPs will be influenced by both the number of available 'truthset' variants and their concordance with assay readouts. Methods: We first reviewed 112 sets of ClinGen gene-specific classification specifications (CSPECs) to assess methodologies they applied for truthset assembly and clinical validation of assays. We then proposed differing rules regarding variant type and stringency of classification by which truthsets might be assembled using ClinVar-extracted classifications. We then examined augmentation of ClinVar-classified truthsets with 'proxy-clinical' benign-classified missense variants systematically assembled applying ACMG/AMP rules (of differing stringencies). In total, these constituted 70 basic approaches to ClinVar-based truthset assembly, which we applied to VHL, BRCA1, BRCA2 and RAD51C. We additionally analysed the impact on the size of the truthsets of changing the specified phenotypes against which ClinVar classification had been submitted. We then applied these truthsets to quantify concordance and allocatable EPs for five large-scale multiplexed functional assays for VHL, BRCA1, BRCA2, and RAD51C. Results The EPs from clinical validation of each assay varied widely according to which truthset was used across 2,120 permutations of gene-truthset-assay combinations. For example, sequentially applying 700 different ClinVar-based truthsets to 2,268 VHL assay variant readouts (70 basic ClinVar-based approaches, augmented by examining 5 different phenotypes for each basic approach, and separate validation against two defined deleterious zones), the evidence strength allocatable for pathogenicity ranged from nil to strong evidence (0.0 to 5.6 EPs); for benignity it ranged from supporting to strong evidence (-1.5 to -6.4 EPs). Clinical validation using truthsets comprising just ClinVar-classified missense variants typically resulted in lower EPs than truthsets comprising protein truncating (PTV) and synonymous variants; this was more due to paucity of ClinVar-classified missense truthset variants than poorer concordance. Augmentation with larger 'proxy-clinical' benign-classified missense truthsets typically improved evidence allocatable for pathogenicity, with improved power negating modest reduction in concordance. Conclusions EPs can be improved by augmentation with systematically-generated 'proxy-clinical' benign-classified missense variants and/or reduction of truthset stringency. Explicit prescriptive clinical guidance is urgently required to improve consistency in clinical validation of functional assays and consequent evidence application for clinical variant classification.
Dulcic, D.; Mandic, K.; Hrsak, D.; Baresic, A.
Show abstract
Common variants detected by the genome-wide association studies (GWAS) create a wealth of knowledge on genetic component of individual traits and diseases. Elucidating the molecular mechanism behind the vast majority of these variants that are found to be non-coding remains a largely unsolved task, especially when distal and pleiotropic interactions between regulatory elements where these variants occur and gene promoters are taken into account. Focusing on four diseases with immune-mediated mechanisms namely ulcerative colitis, Crohn's disease, primary sclerosing cholangitis and ankylosing spondylitis, we demonstrate the utility of the targPred tool, providing prediction of genes targeted by the regulatory variants. We demonstrate that taking into account evolutionary and comparative genomic data, previously unobserved mechanistic trends (the platelet, vascular and sterol clusters) can be detected in terms of implicated genes targeted by the regulatory elements containing common variants, shared between all four diseases, as well as specific trends for subsets of diseases, e.g. two IBD phenotypes. We also elucidate a clinically-relevant target COG6 shared between IBD and PSC, as well as a whole range of other target genes missed by the conventional SNP-to-gene assignments methods.
Viz-Lasheras, S.; Dacosta, A.; Rivero-Calle, I.; Martinon-Torres, F.; EUCLIDS, GENDRES, PERFORM, and DIAMONDS consortia, ; Gomez-Carballa, A.; Salas, A.
Show abstract
Accurate discrimination between viral, bacterial, and inflammatory diseases in febrile children remains a major clinical challenge that contributes to diagnostic uncertainty, inappropriate antimicrobial use, and suboptimal clinical management. Host blood transcriptomics offer a promising strategy to improve diagnostic precision. The present study represents the largest integrative multi-cohort pediatric study of transcriptomic biomarker discovery, validation, and confirmation reported to date, integrating harmonized public transcriptomic datasets with an independent confirmation cohort comprising well-phenotyped patients to identify parsimonious host-response signatures for differentiating viral, bacterial, and inflammatory diseases. Transcriptomic signatures were derived from an integrated retrospective microarray multi-cohort (n=1,683), independently validated in a retrospective RNA-seq cohort (n=767), and confirmed by digital PCR in an independent cohort (n=29), demonstrating reproducibility across patient populations, transcriptomic technologies, and analytical platforms. The analysis identified binary signatures and a unified multiclass classifier that consistently achieved high diagnostic accuracy across all three study phases and outperformed more than 30 published host transcriptomic signatures. Decision curve analysis showed substantially greater clinical net benefit than C-reactive protein across clinically relevant decision thresholds. These findings provide a strong foundation for clinically deployable molecular diagnostics to improve patient triage, antimicrobial stewardship, and precision medicine in childhood infections.
Luijts, T.; Hoogstoel, S.; Pappaert, E.; De Meester, E.; Van Nieuwerburgh, F.; Van Hamme, E.; De Schepper, S.; Willaert, W.; Vral, A.; Hoorens, I.; Van den Eynden, J.
Show abstract
Spatial transcriptomics (ST) has revolutionized our understanding of tumor biology but inherently lacks information on the upstream somatic driver mutations. We developed a spatially-aware graph convolutional neural network (MuT-GCNN) that infers TP53 clones directly from ST data. MuT-GCNN was trained on virtual ST slides with clones simulated from a large collection of existing RNA and matched DNA sequencing data. The model is highly performant with precision and recall values exceeding 95% in most analysed cancer types. It is sensitive for single hit mutations and is primarily informed by the expression of p53 signalling genes in cancer cells. After demonstrating the potential of the model on publicly available squamous cell carcinoma (SCC) data, a direct validation was performed using ST and matched DNA sequencing from serial slices obtained from 4 cutaneous SCC samples. With the increasing availability of ST data and upcoming ST atlases, MuT-GCNN can unveil the location of (sub)clonal alterations in TP53, the most frequently mutated gene in human cancer.
Lee, T. S. E.; Nguyen, L.; Forde, B. M.; Maidment, T.; Ye, S.; Henderson, A.; Playford, E. G.; Runnegar, N.; Henderson, B.; Watson, C.; Lindsay, M.; Bursle, E.; Douglas, J.; Hume, J.; Paterson, D. L.; Kidd, T.; Graves, B.; Hume, A.; Hall, M. B.; Schembri, M. A.; Beatson, S. A.; Harris, P. N. A.; Roberts, L. W.
Show abstract
OXA-48-like carbapenemases have been historically rare, however steady increases both locally and globally have warranted further investigation into their spread. Here we present the largest genomic analysis of blaOXA-181-producing bacteria in Australia to date, focusing on a single jurisdiction over seven years (2017 -- 2024). The initial investigation was prompted by an outbreak in 2017, where enhanced genomic surveillance in a single hospital identified 85 outbreak isolates related to an imported Escherichia coli ST38, carrying blaOXA-181 on an IncX3/colKP3 plasmid (previously reported as pOXA181). After four months of intensive infection control, the initial outbreak strain was eliminated. To confirm the outbreak plasmid was also contained, we collected all blaOXA-181-positive isolates from the same jurisdiction over subsequent years and sequenced with both Illumina and Oxford Nanopore Technologies to investigate clonal and mobile genetic element mediated spread. While continued surveillance post-2017 did not identify the same E. coli strain following the outbreak, pOXA181 plasmids were identified in >70% of surveillance isolates, with minimal genetic changes, which initially suggested local plasmid-mediated spread. Additional comparison to a global collection of pOXA181 plasmids found that epidemiologically unrelated pOXA181 plasmids were near identical, with no rearrangements and low, or no, single nucleotide polymorphisms. This suggests the mutation rate of pOXA-181 is incompatible with recent genomic transmission inference. This study highlights the current genomic epidemiology and drivers of blaOXA-181 and further demonstrates the necessity for detailed understanding of plasmid evolutionary rates to inform genomic surveillance.